Semantic Web • Week 02
XML & XML Schema
Structured data representation, validation with XSD, and the XML model of the allergy data
Graduate Semantic Web Course • CMPE 583
Week 02 • Objectives
By the end of this week
- you will be able to tell well-formed XML from valid XML in practice;
- you will be able to model the allergy data in XML and justify your element/attribute decisions;
- you will be able to write datatype, facet, cardinality and key constraints in XSD;
- you will be able to validate an XML document against a schema in Java;
- you will be able to explain why the tree model of XML is not enough for the graph model of RDF.
Recap from last week
- The Document Web encodes presentation; the Data Web encodes meaning.
- The stack: IRI → XML → RDF → RDFS → OWL → rules/queries.
- The risk chain in the project: Person → Product → FoodAdditives → Allergy.
- This week we are on the second step of the stack: syntax and validation.
Output of this week
products.xml
persons.xml
allergy.xsd
Validate.java
01
Solution of Assignment 1
Toolchain setup, ontology inventory and the class/individual distinction.
Solution 1.1 — Checking the environment
$ java -version
openjdk version "1.8.0_392"
$ mvn -v
Apache Maven 3.9.6
Protégé 5.6.4
Tabs: Entities · Individuals
SWRLTab · OntoGraf
| Check | Expected evidence |
| Protégé starts | Version in the title bar |
| SWRLTab is visible | Screenshot of the tab |
| The OWL file opens | The class tree is populated |
| JDK / Maven | Version output |
If SWRLTab is not visible: File → Check for plugins → install the SWRLTab and SWRLAPI plugins and restart Protégé.
Solution 1.2 — Inventory of ALLERGY_FIXED.owl
| Entity type | Count | Examples |
| Class | 6 | Person, Product, FoodAdditives, Allergy, Adult, PersonAtRisk |
| Object property | 9 | Contain, Triggers, hasAllergy, ChooseProduct, Effected_Allergen + 4 sub-properties |
| Data property | 6 | hasName, hasAge, hasWeight, hasHeight, hasBMI, hasProductName |
| Individual | 17 | 4 persons, 4 products, 5 additives, 4 allergies |
| SWRL rule | 7 | S1 … S7 |
| Disjointness axiom | 1 | AllDisjointClasses (4 top-level classes) |
The classes Adult and PersonAtRisk have no asserted members; membership comes from the rules.
Solution 1.3 — Class or individual?
| Term | Correct modelling | Reason |
| Food additive | Class (FoodAdditives) | A kind, it has members |
| Nisin | Individual | One specific substance |
| Lactose allergy | Individual (Lactose) | A specific member of the Allergy class |
| Product | Class | It covers barcoded individuals |
| Eti Chocolate | Individual + hasProductName | The barcode is identity, the name is data |
| Person at risk | Class, populated by a rule | Membership is computed, not asserted |
Solution 1.4 — Modelling EAN_00005
EAN_00005 a Product ;
Contain Whey_Protein , Wheat_Starch ;
hasProductName "Protein Biscuit" .
Whey_Protein a FoodAdditives ;
Triggers Lactose .
Wheat_Starch a FoodAdditives ;
Triggers Gluten .
TC_004 (allergic to Egg and Gluten) chooses this product:
S6 → Effected_Allergen(TC_004, Wheat_Starch)
S7 → PersonAtRisk(TC_004)
Whey_Protein triggers nothing here: TC_004 has no lactose allergy, so hasAllergy(?p, ?al) in the rule body does not bind.
02
XML Fundamentals
The tree model, well-formedness rules, namespaces.
What is XML?
- A text format that marks data up with tags and is independent of any application.
- There is no fixed tag set; the domain expert defines the tags.
- It carries structure, not meaning — the meaning stays in the schema and in the application.
<?xml version="1.0" encoding="UTF-8"?>
<product ean="EAN_00004">
<name>Eti Chocolate</name>
<additives>
<additive>Nisin</additive>
</additives>
</product>
Conditions for being well-formed
- There must be exactly one root element.
- Every opened tag must be closed.
- Nesting must be in the correct order.
- Tag names are case sensitive.
- Attribute values must be quoted.
<!-- HATALI -->
<product ean=EAN_00004> ← no quotes
<Name>Eti</name> ← case mismatch
<additives><additive>Nisin
</additives></additive> ← wrong nesting order
</product>
No parser reads a document that is not well-formed — it is the minimum condition before validation.
The parts of an XML document
<?xml version="1.0" encoding="UTF-8"?> ← prolog
<!-- Product catalogue of the allergy project --> ← comment
<products count="4"> ← root element + attribute
<product ean="EAN_00003"> ← child element
<name>Dardanel Ton</name> ← text (PCDATA)
<additive ref="Casein"/> ← empty element
</product>
</products>
A document is a tree: every node has exactly one parent. This restriction is why we move to RDF in Week 03.
Element or attribute?
Attribute-heavy
<person tc="TC_001" name="Ayse"
age="38" weight="67.5"
height="1.68"/>
Short; but it cannot repeat and cannot carry structure.
Element-heavy
<person tc="TC_001">
<name>Ayse</name>
<age>38</age>
<allergy ref="Lactose"/>
<allergy ref="Fish"/>
</person>
multiple values and extension are not possible.
Rule of thumb: identity and metadata as attributes, field data as elements. A person may have more than one allergy, so allergy must be an element.
Namespaces: same name, different meaning
<cat:products
xmlns:cat="http://EMU/catalog#"
xmlns:med="http://EMU/medical#">
<cat:product ean="EAN_00003">
<cat:name>Dardanel Ton</cat:name>
<med:risk level="high"/>
</cat:product>
</cat:products>
- xmlns:prefix="IRI" declarations resolve name clashes.
- It is the IRI that carries the meaning, not the prefix.
- A declaration without a prefix (xmlns=) sets the default namespace.
- The OWL file of the project uses the same mechanism: rdf:, owl:, swrl:.
The XML header of the project
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
xmlns:xsd="http://www.w3.org/2001/XMLSchema#"
xmlns:rdfs="http://www.w3.org/2000/01/rdf-schema#"
xmlns:owl="http://www.w3.org/2002/07/owl#"
xml:base="http://EMU/AllergyOntology"
xmlns="http://EMU/AllergyOntology#"
xmlns:swrl="http://www.w3.org/2003/11/swrl#">
| Prefix | Role |
| rdf, rdfs | Graph and vocabulary syntax |
| owl | Ontology constructs |
| xsd | Datatypes (int, double, string) |
| swrl | Rule axioms |
| (default) | The project's own terms: Person, Nisin … |
Special characters and CDATA
| Character | Entity |
| < | < |
| > | > |
| & | & |
| " | " |
| ' | ' |
<ingredients>
Contains milk & cocoa
</ingredients>
<label><![CDATA[
E322 (soya lesitini) < 0.5%
]]></label>
Percent signs and "&" are common on food labels; this is the most frequent source of parsing errors.
Encoding: pitfalls with non-ASCII characters
<?xml version="1.0" encoding="UTF-8"?>
<product ean="EAN_00006">
<name>Ulker Chocolate Wafer</name>
<note>Contains milk · 30% cocoa</note>
</product>
<!-- if saved as ISO-8859-9: -->
<!-- Ülker Çikolatalı -->
| Pitfall | Result |
| Declared UTF-8, file saved as ANSI | Parsing error |
| UTF-8 with BOM | "Content is not allowed in prolog" |
| Dotless-i case mapping | Use Locale.ROOT in Java |
| Non-ASCII letter in an IRI | Keep local names in ASCII |
Project rule: label text may be in any language, but the local name of an IRI is always ASCII (Soy_Lecitin, label: "Soy lecithin").
The tree model of the document
products
├── product @ean=EAN_00003
│ ├── name "Dardanel Ton"
│ ├── additive @ref=Casein
│ └── additive @ref=Sodium_Ascorbite
└── product @ean=EAN_00004
├── name "Eti Chocolate"
├── additive @ref=Nisin
└── additive @ref=Soy_Lecitin
This tree does not say which allergy an additive triggers; we point to another document with ref — the link is made by the application.
Example: products.xml
<?xml version="1.0" encoding="UTF-8"?>
<products xmlns="http://EMU/allergy/catalog"
count="4">
<product ean="EAN_00001">
<name>ETI Cracker</name>
<additive ref="Alginic_Acid"/>
</product>
<product ean="EAN_00002">
<name>Ulker Damak</name>
<additive ref="Phospore"/>
<additive ref="Soy_Lecitin"/>
</product>
<product ean="EAN_00003">
<name>Dardanel Ton</name>
<additive ref="Casein"/>
<additive ref="Sodium_Ascorbite"/>
</product>
<product ean="EAN_00004">
<name>Eti Chocolate</name>
<additive ref="Ascorbic_Acid"/>
<additive ref="Nisin"/>
<additive ref="Soy_Lecitin"/>
</product>
</products>
Additives point to a separate document with ref ; there is no direct link between a product and an allergy.
Example: additives.xml
<additives xmlns="http://EMU/allergy/catalog">
<additive id="Nisin" code="E234">
<label>Nisin</label>
<triggers allergy="Lactose"/>
</additive>
<additive id="Casein" code="E290">
<label>Casein</label>
<triggers allergy="Lactose"/>
</additive>
<additive id="Soy_Lecitin" code="E322">
<label>Soy Lecithin</label>
<triggers allergy="Egg"/>
</additive>
<additive id="Ascorbic_Acid" code="E300">
<label>Ascorbic Acid</label>
</additive>
</additives>
Ascorbic_Acid has no triggers — exactly the situation in the ontology.
In XML this "gap" is just a missing value. In OWL, under the Open World Assumption, it means "unknown" — the same data, two different readings.
Example: persons.xml
<persons xmlns="http://EMU/allergy/profile">
<person tc="TC_001">
<name>Ayse</name>
<age>38</age>
<weight unit="kg">67.5</weight>
<height unit="m">1.68</height>
<allergy ref="Lactose"/>
<choice ean="EAN_00004"/>
</person>
<person tc="TC_003">
<name>MEHMET</name>
<age>35</age>
<weight unit="kg">93.0</weight>
<height unit="m">1.87</height>
<allergy ref="Fish"/>
<allergy ref="Lactose"/>
<choice ean="EAN_00003"/>
</person>
</persons>
BMI is absent here: computed values are not kept in the source data — rule S4 will produce it in Week 07.
03
Validation: DTD and XML Schema
Datatypes, constraints, keys and reading error messages.
The difference between well-formed and valid
| Aspect | Well-formed | Valid |
| What it checks | Syntax | Conformance to the schema |
| Reference | The XML 1.0 rules | The DTD / XSD document |
| "age=abc" | Valid | Error: int expected |
| Unknown element | Valid | Error |
| Required field missing | Valid | Error |
Label data comes from an external source, so validation is compulsory: dirty data in the ontology produces wrong risk inferences.
Validation with a DTD, and its limits
<!ELEMENT products (product+)>
<!ELEMENT product (name, additive*)>
<!ATTLIST product ean ID #REQUIRED>
<!ELEMENT name (#PCDATA)>
<!ELEMENT additive EMPTY>
<!ATTLIST additive ref IDREF #REQUIRED>
- No datatypes: age may be "abc".
- No namespace support.
- Numeric ranges, patterns and decimal constraints cannot be expressed.
- It is not XML itself, so tooling is harder.
This is why we use XSD in the project; we learn DTD only to read legacy documents.
The skeleton of an XML Schema document
<?xml version="1.0" encoding="UTF-8"?>
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"
xmlns="http://EMU/allergy/catalog"
targetNamespace="http://EMU/allergy/catalog"
elementFormDefault="qualified">
<xs:element name="products" type="ProductsType"/>
<!-- type definitions go here -->
</xs:schema>
targetNamespace states which namespace the schema defines; the document must use the same namespace.
Built-in datatypes
| XSD type | Example value | Field in the project |
| xs:string | "Eti Chocolate" | hasName, hasProductName |
| xs:int | 38 | hasAge |
| xs:double | 67.5 | hasWeight, hasHeight, hasBMI |
| xs:boolean | true | label validated? |
| xs:date | 2026-09-08 | expiry date |
| xs:ID / xs:IDREF | EAN_00004 | barcode and reference |
The same type names appear on the OWL side: rdfs:range = &xsd;double. XSD datatypes are the shared vocabulary of the two worlds.
Define your own type: facets
<xs:simpleType name="AgeType">
<xs:restriction base="xs:int">
<xs:minInclusive value="0"/>
<xs:maxInclusive value="120"/>
</xs:restriction>
</xs:simpleType>
<xs:simpleType name="WeightType">
<xs:restriction base="xs:double">
<xs:minExclusive value="0"/>
<xs:maxInclusive value="400"/>
</xs:restriction>
</xs:simpleType>
| Facet | What it does |
| minInclusive | Lower bound (inclusive) |
| maxExclusive | Upper bound (exclusive) |
| length | Length |
| pattern | Regular expression |
| enumeration | Set of allowed values |
| fractionDigits | Decimal digits |
Enumeration: allergy types
<xs:simpleType name="AllergyType">
<xs:restriction base="xs:string">
<xs:enumeration value="Lactose"/>
<xs:enumeration value="Gluten"/>
<xs:enumeration value="Egg"/>
<xs:enumeration value="Fish"/>
</xs:restriction>
</xs:simpleType>
Here XSD builds a closed world: a value outside the list is an error.
The OWL counterpart owl:oneOf can express this; in the project, however, we keep allergies as individuals, because adding a new allergy type should not force a schema change.
Pattern: barcode and identifier format
<xs:simpleType name="EanType">
<xs:restriction base="xs:string">
<xs:pattern value="EAN_[0-9]{5}"/>
</xs:restriction>
</xs:simpleType>
<xs:simpleType name="TcType">
<xs:restriction base="xs:string">
<xs:pattern value="TC_[0-9]{3}"/>
</xs:restriction>
</xs:simpleType>
| Value | EanType result |
| EAN_00004 | Valid |
| EAN_4 | Error — five digits required |
| ean_00004 | Error — upper case required |
Complex type: product
<xs:complexType name="ProductType">
<xs:sequence>
<xs:element name="name" type="xs:string"/>
<xs:element name="additive" type="AdditiveRefType"
minOccurs="0" maxOccurs="unbounded"/>
</xs:sequence>
<xs:attribute name="ean" type="EanType" use="required"/>
</xs:complexType>
<xs:complexType name="AdditiveRefType">
<xs:attribute name="ref" type="xs:string" use="required"/>
</xs:complexType>
minOccurs="0": a product with no declared additive is still valid — a missing value does not break the schema.
sequence, choice, all
| Construct | Meaning | Allergy example |
| sequence | All of them, in the given order | name, then the additive list |
| choice | Only one of them | ya ean ya internalCode as identifier |
| all | All of them, order free | weight, height, age |
| group | A reusable group | the measurement block |
<xs:choice>
<xs:element name="ean" type="EanType"/>
<xs:element name="internalCode" type="xs:string"/>
</xs:choice>
Cardinality constraints
| Declaration | Meaning |
| minOccurs="1" | Required (default) |
| minOccurs="0" | Optional |
| maxOccurs="unbounded" | Unbounded repetition |
| use="required" | Required attribute |
The OWL counterpart
Product ⊑ ≥1 Contain.FoodAdditives
An XSD constraint rejects data; an OWL restriction produces an inference. The same sentence, two different behaviours — covered in detail in Week 05.
key and keyref: referential integrity
<xs:element name="catalog" type="CatalogType">
<xs:key name="additiveKey">
<xs:selector xpath="additives/additive"/>
<xs:field xpath="@id"/>
</xs:key>
<xs:keyref name="additiveRef" refer="additiveKey">
<xs:selector xpath="products/product/additive"/>
<xs:field xpath="@ref"/>
</xs:keyref>
</xs:element>
This way a typo such as <additive ref="Nisiin"/> is caught during validation — a wrong individual never reaches the ontology.
The full schema: allergy.xsd
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"
targetNamespace="http://EMU/allergy/catalog"
xmlns="http://EMU/allergy/catalog"
elementFormDefault="qualified">
<xs:element name="products">
<xs:complexType>
<xs:sequence>
<xs:element name="product" type="ProductType"
maxOccurs="unbounded"/>
</xs:sequence>
<xs:attribute name="count" type="xs:int"/>
</xs:complexType>
</xs:element>
<xs:complexType name="ProductType">
<xs:sequence>
<xs:element name="name" type="xs:string"/>
<xs:element name="additive"
type="AdditiveRefType" minOccurs="0"
maxOccurs="unbounded"/>
</xs:sequence>
<xs:attribute name="ean" type="EanType"
use="required"/>
</xs:complexType>
<xs:simpleType name="EanType">
<xs:restriction base="xs:string">
<xs:pattern value="EAN_[0-9]{5}"/>
</xs:restriction>
</xs:simpleType>
</xs:schema>
Reading validation errors
<product ean="EAN_4">
cvc-pattern-valid: 'EAN_4' is not facet-valid with respect to pattern 'EAN_[0-9]{5}'
<age>abc</age>
cvc-datatype-valid.1.2.1: 'abc' is not a valid value for 'int'
<product> <!-- no ean -->
cvc-complex-type.4: Attribute 'ean' must appear on element 'product'
The cvc-* code at the start of the message tells you which rule was violated; it is the fastest way to debug.
Validation in Java: Validate.java
import javax.xml.XMLConstants;
import javax.xml.validation.*;
import org.xml.sax.SAXException;
import java.io.File;
public class Validate {
public static void main(String[] a) throws Exception {
SchemaFactory sf = SchemaFactory.newInstance(
XMLConstants.W3C_XML_SCHEMA_NS_URI);
Schema schema = sf.newSchema(new File("allergy.xsd"));
Validator v = schema.newValidator();
try {
v.validate(new javax.xml.transform.stream.StreamSource(
new File("products.xml")));
System.out.println("products.xml GECERLI");
} catch (SAXException e) {
System.out.println("HATA: " + e.getMessage());
}
}
}
Choosing a parser: DOM, SAX, StAX
| Model | Approach | Memory | In the allergy project |
| DOM | Loads the document as a tree | High | Small catalogue, navigation with XPath |
| SAX | Event based, single pass | Very low | A store dump with thousands of products |
| StAX | Pull-based streaming | Low | Partial reading, early exit |
| JAXB | Object mapping | Medium | Binding to the Product/Person classes |
Rule of thumb: DOM if the data fits in memory and you need random access; SAX or StAX if large data arrives as a stream.
Java + DOM: reading the catalogue
DocumentBuilderFactory f =
DocumentBuilderFactory.newInstance();
f.setNamespaceAware(true);
Document doc = f.newDocumentBuilder()
.parse(new File("products.xml"));
NodeList ps = doc.getElementsByTagNameNS(
"http://EMU/allergy/catalog", "product");
for (int i = 0; i < ps.getLength(); i++) {
Element p = (Element) ps.item(i);
String ean = p.getAttribute("ean");
NodeList as = p.getElementsByTagNameNS("*","additive");
for (int j = 0; j < as.getLength(); j++)
System.out.println(ean + " Contain " +
((Element) as.item(j)).getAttribute("ref"));
}
EAN_00001 Contain Alginic_Acid
EAN_00002 Contain Phospore
EAN_00002 Contain Soy_Lecitin
EAN_00003 Contain Casein
EAN_00003 Contain Sodium_Ascorbite
EAN_00004 Contain Ascorbic_Acid
EAN_00004 Contain Nisin
EAN_00004 Contain Soy_Lecitin
This output is already in triple form — in Week 08 the same lines become axioms through the OWL API.
Java + SAX: reading as a stream
SAXParserFactory.newInstance().newSAXParser().parse(
new File("market_dump.xml"),
new DefaultHandler() {
String ean;
public void startElement(String u, String l, String q, Attributes at) {
if ("product".equals(l)) ean = at.getValue("ean");
if ("additive".equals(l)) risk(ean, at.getValue("ref"));
}
});
// risk(): if the additive triggers Lactose, add it to the warning list
SAX keeps nothing in memory: a 500 MB store dump can be scanned with constant memory; in exchange you cannot go back and navigate.
Selecting data with XPath
| XPath | Result |
| /products/product/@ean | All barcodes |
| //product[additive/@ref='Nisin']/name | Names of products containing Nisin |
| count(//product[@ean='EAN_00004']/additive) | 3 |
| //person[age>=18]/name | Names of adults |
| //person[allergy/@ref='Lactose']/@tc | TC_001, TC_002, TC_003 |
Note: the last query finds only asserted allergies. The question "who is at risk" cannot be answered in XPath, because the additive → allergy chain lies outside the document.
A report with XQuery: which product affects whom?
for $p in doc("persons.xml")//person
let $ean := $p/choice/@ean
let $prod := doc("products.xml")
//product[@ean = $ean]
for $a in $prod/additive/@ref
let $trg := doc("additives.xml")
//additive[@id = $a]/triggers/@allergy
where $trg = $p/allergy/@ref
return <risk tc="{$p/@tc}" ean="{$ean}"
additive="{$a}"/>
<risk tc="TC_001" ean="EAN_00004"
additive="Nisin"/>
<risk tc="TC_002" ean="EAN_00003"
additive="Casein"/>
<risk tc="TC_003" ean="EAN_00003"
additive="Casein"/>
<risk tc="TC_003" ean="EAN_00003"
additive="Sodium_Ascorbite"/>
We will write the same result in a single SWRL rule (S6). The difference: here you build the chain; there the engine applies the rule and writes the result permanently into the ontology.
From XML to RDF/XML with XSLT
<xsl:template match="product">
<owl:NamedIndividual
rdf:about="#{@ean}">
<rdf:type rdf:resource="#Product"/>
<hasProductName>
<xsl:value-of select="name"/>
</hasProductName>
<xsl:for-each select="additive">
<Contain rdf:resource="#{@ref}"/>
</xsl:for-each>
</owl:NamedIndividual>
</xsl:template>
<!-- output -->
<owl:NamedIndividual rdf:about="#EAN_00004">
<rdf:type rdf:resource="#Product"/>
<hasProductName>Eti Chocolate
</hasProductName>
<Contain rdf:resource="#Ascorbic_Acid"/>
<Contain rdf:resource="#Nisin"/>
<Contain rdf:resource="#Soy_Lecitin"/>
</owl:NamedIndividual>
Using this transformation instead of typing label data into the ontology by hand is the most practical way to feed the project with real data.
04
Is XML Enough?
The difference between a tree and a graph, and why we move to RDF.
The tree model and the graph model
XML: a tree
person
└── allergy @ref="Lactose"
(the link is only a name)
The reference is text; only the application code knows its meaning.
RDF: a graph
TC_001 hasAllergy Lactose .
Nisin Triggers Lactose .
EAN_00004 Contain Nisin .
Nodes are shared; inference runs along the chain.
Why is XML alone not enough?
| What is missing | Consequence | Layer that solves it |
| Order carries meaning | Same information, different tree | RDF (unordered triples) |
| Links are unnamed | Meaning lives in the code | RDF predicates |
| No class / subclass | Hierarchy in the code | RDFS |
| No constraints or logic | Contradictions cannot be found | OWL |
| No inference | Only what is written is known | Reasoner + SWRL |
| No global identity | Data merging by hand | IRI |
XML remains indispensable: our RDF/XML and OWL files are XML — it stays as the transport layer.
The same product, two representations
products.xml
<product ean="EAN_00003">
<name>Dardanel Ton</name>
<additive ref="Casein"/>
<additive ref="Sodium_Ascorbite"/>
</product>
ALLERGY_FIXED.owl
<owl:NamedIndividual rdf:about="#EAN_00003">
<rdf:type rdf:resource="#Product"/>
<Contain rdf:resource="#Casein"/>
<Contain rdf:resource="#Sodium_Ascorbite"/>
<hasProductName rdf:datatype="&xsd;string">
Dardanel Ton</hasProductName>
</owl:NamedIndividual>
The syntax is almost identical; the difference is the globally identified link built with rdf:resource. That difference is exactly the topic of Week 03.
Checklist for catalogue data
- Every document declares a namespace; use the prefix consistently.
- Identifiers are ASCII and pattern constrained (EAN_[0-9]{5}).
- Carry the unit of measure as an attribute (unit="kg").
- Do not keep a computed value (BMI) in the source data.
- Keep the additive → allergy mapping in exactly one place.
- Validate with XSD before every load; log the error.
- Save as UTF-8 without a BOM.
- Keep the transformation (XSLT) under version control; never fix output by hand.
This list is applied directly in Week 08, when the ontology is fed with real catalogue data.
05
Assignment and Project Step
Build your own product catalogue with XML + XSD.
Assignment 2 — Catalogue XML and its schema
- Write products.xml for five packaged products of your own choice.
- Define the additive → allergy mapping in additives.xml.
- Write allergy.xsd with a barcode pattern, an age range and an allergy enumeration.
- Enforce referential integrity with keyref.
- Produce two invalid documents and report the validation messages.
Deliverable
Four files + a 2-page report: design decisions (element vs attribute), the reasons for your facets, and an interpretation of the error messages.
We will discuss the solution at the start of Week 03.
Assessment criteria
| Criterion | Weight | Expected |
| Well-formed + valid documents | 25% | Validation passes without errors |
| Expressiveness of the schema | 30% | Facets, patterns, enumeration and cardinality are used |
| Referential integrity | 20% | key/keyref works |
| Design rationale | 15% | Element/attribute decisions are defended |
| Error analysis | 10% | cvc codes are interpreted correctly |
References
- W3C — Extensible Markup Language (XML) 1.0, 5th Edition.
- W3C — XML Schema Part 0: Primer; Part 1: Structures; Part 2: Datatypes.
- W3C — Namespaces in XML 1.0; XML Path Language (XPath) 3.1.
- Harold, E. R., Means, W. S. — XML in a Nutshell, O'Reilly.
- Allemang & Hendler — Semantic Web for the Working Ontologist, Part 3 (moving to RDF).
Summary • 1 / 2
XML and validation
- XML carries structure, not meaning; the meaning stays in the schema and the application.
- Well-formed is the minimum condition; valid means conforming to the schema.
- XSD gives datatypes, facets, cardinality and key constraints; DTD does not.
- Identity and metadata become attributes; repeatable field data becomes elements.
- Validation is the first line of defence against dirty data entering the ontology.
Summary • 2 / 2
Contribution to the project and the next step
- The allergy data was split into three documents: products, additives, persons.
- Barcode and identifier formats were secured with patterns.
- The skeleton of the XML → RDF/XML transformation was built with XSLT.
- XPath queries asserted data; it cannot query the risk chain.
In Week 03
RDF & RDFS: the triple model, thinking in graphs, Turtle syntax and building the allergy vocabulary.
Also: the detailed solution of Assignment 2.
Review Questions
Test yourself
- Can a document be well-formed but not valid? Give an example.
- Why did we model the allergy list as elements rather than attributes?
- minOccurs="0" and the ≥1 restriction in OWL — what is the behavioural difference?
- Which error does keyref catch, and which one does it not?
- Why can the "persons at risk" query not be written in XPath?
- Explain the relation between a namespace prefix and an IRI.
- What is the difference in world assumption between XSD enumeration and OWL oneOf?
- In the XSLT transformation, why was rdf:resource used instead of a text value?
Exercise • In class
Test the schema against an invalid document
The document below contains three validation errors. Find them, predict the cvc message and fix them.
<products count="two">
<product ean="EAN_105">
<additive ref="Nisin"/>
<name>Test Bar</name>
</product>
</products>
Hint
Think in order: attribute type, pattern conformance, element order inside xs:sequence.
The solution comes in the assignment-solution part of Week 03.